Papers with text-based models

12 papers
Augmenting pre-trained language models with audio feature embedding for argumentation mining in political debates (2023.findings-eacl)

Copied to clipboard

Challenge: Existing studies on the integration of multimodality with text and audio in natural language processing tasks have focused on the use of image and text for emotion recognition, fake news detection and document image classification.
Approach: They propose to integrate audio features with text in a task of argumentation mining using a previously reported dataset and an audio-enhanced version.
Outcome: The proposed model outperforms text-based models on a dataset of 28,850 utterances on 'argumentation mining' with limited data.
Fine-Tuning Pre-Trained Language Models with Gaze Supervision (2024.acl-short)

Copied to clipboard

Challenge: Existing pre-trained language models lack a gaze module to exploit cognitive signals.
Approach: They propose to integrate a gaze module into pre-trained language models at the fine-tuning stage to exploit cognitive signals.
Outcome: The proposed model improves performance on the GLUE benchmark and standard fine-tuning and text augmentation baselines.
A Word is Worth A Thousand Dollars: Adversarial Attack on Tweets Fools Stock Prediction (2022.naacl-main)

Copied to clipboard

Challenge: Existing models are vulnerable to adversarial attacks, but their vulnerability is underexplored.
Approach: They propose to concatenate a perturbed but semantically similar tweet into a model that fools stock prediction models.
Outcome: The proposed method achieves consistent success rates and causes significant monetary loss in trading simulation by simply concatenating a perturbed but semantically similar tweet.
Deja vu: Contrastive Historical Modeling with Prefix-tuning for Temporal Knowledge Graph Reasoning (2024.findings-naacl)

Copied to clipboard

Challenge: Existing text-based methods for Temporal Knowledge Graph Reasoning struggle to balance textual knowledge and temporal information with expensive purpose-built training strategies.
Approach: They propose a Contrastive historical modeling framework with prefix-tuning for TEmporal Reasoning that feeds history-contextualized text into the pseudo-Siamese encoders to strike a textual-temporal balance.
Outcome: The proposed framework achieves superior performance on four transductive and three few-shot inductive TKGR benchmarks.
Multimodal Extraction and Recognition of Arabic Implicit Discourse Relations (2025.coling-main)

Copied to clipboard

Challenge: Identifying implicit discourse relations in written text is challenging, but it is also crucial to understand them in spoken discourse.
Approach: They propose a method for implicit discourse relation identification that uses audio and text data to extract semantically equivalent pairs of implicit and explicit discourse markers.
Outcome: The proposed method outperforms audio-based models but can be augmented by combining text and audio features.
Beyond Transcripts: A Renewed Perspective on Audio Chaptering (2026.acl-long)

Copied to clipboard

Challenge: despite its relevance, research on audio chaptering remains limited and predominantly textbased . authors: audio chapterers can't be used linearly because they skim, scrub timelines, jump to relevant moments . acoustic features and learning representations are not used for audio chapterer evaluation .
Approach: They propose to use audio-only architecture to automatically segment audio into coherent sections . they compare audio-based models with acoustic features and a novel audio-oriented architecture .
Outcome: The proposed audio-only architecture outperforms text-based approaches on acoustic features and LLMs.
Speech language models lack important brain-relevant semantics (2024.acl-long)

Copied to clipboard

Challenge: Recent work shows that text-based language models predict both text- and speech-evoked brain activity.
Approach: They remove low-level stimulus features from language models to assess their impact on alignment with fMRI brain recordings during reading and listening.
Outcome: The proposed model removes low-level features from fMRI brain recordings to assess their impact on alignment with fmr recordings.
Understanding the Language of Political Agreement and Disagreement in Legislative Texts (2020.acl-main)

Copied to clipboard

Challenge: Despite the fact that state-level legislation is rarely discussed, it has a dramatic influence on the everyday life of residents of the respective states.
Approach: They propose a large-scale dataset linking state bills and legislator information, geographical information about their districts, and donations and donors’ information.
Outcome: The proposed model improves over strong text-based models by integrating the state-level text and the legislative context.
MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that multimodal, multimodal approaches to lyrics translation are more effective than text-only approaches.
Approach: They propose a multilingual, multimodal benchmark for singable lyrics translation . they propose syllable-constrained audio-video LLM with Chain-of-Thought .
Outcome: The proposed system outperforms text-based models in singability and contextual accuracy.
SPIRIT: Patching Speech Language Models against Jailbreak Attacks (2025.emnlp-main)

Copied to clipboard

Challenge: Speech language models (SLMs) enable natural interactions via spoken instructions, which more effectively capture user intent by detecting nuances in speech.
Approach: They propose post-hoc patching defenses to intervene during inference by modifying the SLM’s activations that improve robustness up to 99% with negligible impact on utility and without any re-training.
Outcome: The proposed defenses improve robustness up to 99% with negligible impact on utility and (ii) without any re-training.
Neighboring Words Affect Human Interpretation of Saliency Explanations (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies found that superficial factors such as word length can distort human interpretation of the communicated saliency scores.
Approach: They conduct a user study to examine how the marking of a word’s *neighboring words* affect the explainee’s perception of the word’ s importance in the context of . a saliency explanation.
Outcome: The findings question whether text-based saliency explanations should continue to be communicated at word level and inform future research on alternative methods.
Colorful Talks with Graphs: Human-Interpretable Graph Encodings for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Graph problems require reasoning over explicit structure, permutation invariance, and computationally complex relationships, creating a mismatch with the representations of text-based models.
Approach: They propose a human-interpretable structural encoding strategy that injects graph structure directly into natural language prompts.
Outcome: The proposed method improves performance on synthetic and real-world datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations